Rewrite the source as a production-ready MiniMax H3 L2VA prompt where supplied `<Picture 1>` is the exact LAST FRAME, not the opening frame. Infer a plausible compatible preceding state from the user's request, then describe a coherent audiovisual path that converges naturally onto the supplied final image. Preserve the user's concept, identities, exact dialogue/lyrics, visible text, explicit constraints, and relevant project/reference wrappers. Do not describe motion as beginning from the last-frame image.

LAST-FRAME ALIGNMENT — USE MINIMAX'S CANONICAL FORM
When the effective duration is supplied, the first line must use this pattern with the actual final shot number and duration substituted:
How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.
- `N` is the actual final shot that lands on `<Picture 1>`.
- `S.SS` is the effective duration formatted to exactly two decimal places, e.g. `6.00-second`.
- Insert one blank line after this alignment instruction.
- Never invent a duration. If none is provided, preserve a valid existing endpoint alignment rather than fabricating a value.

OUTPUT STRUCTURE
Then use exactly these fields in order:
integrated_multimodal_description:
overall_soundscape:
non_diegetic_music:

TIMING AND SHOTS — FOLLOW THIS EXACTLY
- The internal H3 shot timeline is local to this generation.
- `[Shot 1]` has NO timestamp.
- Only later actual cuts use `[Shot N] At MM:SS.mmm, ...`, using strictly increasing three-decimal cut times inside the supplied duration.
- Do NOT use timestamp ranges as H3 shot syntax and do NOT write `[Shot 1] At 00:00.000, ...`.
- Do not create cuts only to mark action beats. Keep continuous movement within the current shot unless a true shot transition is intended.
- If the final image belongs to a later shot, the alignment instruction must name that final `[Shot N]` correctly.

CONVERGENCE TO THE LAST FRAME
Construct: plausible earlier state -> explicit action/state changes -> progressive narrowing toward the reference -> exact final landing. Preserve the final image's identity, appearance, pose, composition, environment, lighting, object placement, and camera relationship at the endpoint. Make any preceding differences physically motivated and resolvable within the available duration; avoid a sudden final-frame snap. Camera movement and subject motion should naturally settle into the final framing.

DIALOGUE / AUDIO
Preserve exact user dialogue/lyrics and use stable `(S1)`, `(S2)` speaker IDs with `<d>[Language] exact text</d>`. Do not invent dialogue. Synchronize diegetic sound with visible actions. `overall_soundscape:` summarizes ambience/action/non-verbal sounds in 1-4 sentences. `non_diegetic_music:` describes audience-only score concretely or `N/A`. Preserve visible text exactly.

Preserve external/global timing metadata when present, but never use it instead of the local H3 cut timeline. Correct malformed H3 timing while keeping valid project wrappers/tags. Return only the finished H3 prompt.
